rdmsm4x - build - routing intelligence - codex - TASK2026092764 - rdmodelrouter - 20260927-1717
BUILD-RESULT — rdmodelrouter routing intelligence v3
2026-09-27 16:42:04 EDT → 2026-09-27 17:17:55 EDT · rdmsm4x. Builder: codex@rdmsm4x/rmrnight; lead: rdmodelrouter@rdmsm4x/rmr0926. TASK-20260927-64 (parent FEAT-20260926-11). Prototype. Commit-only scope. Branch rmr/routing-intel-20260927; base a32a5df7e3f731087d0406441d150052fd21c106. Worktree /Users/richh/dev/_worktrees/rdmodelrouter-routing-20260927.
Delivered
- Version 0.3.0; default policy v3, preserving explicit v1/v2 configuration/contracts.
- 82 dated model/host catalog entries with per-claim source/as-of/verified metadata: Codex 7, Claude 3 aliases, agy 14, Grok 4, Copilot auto 1, configured Hermes 1, Ollama 33 across six Macs, paid OpenRouter 3, ToshLLM installed GGUFs 16. An embedding-only model and unserved/unverified runtimes are visibly unavailable.
- Deterministic feature/class scorer with evidence, scope/risk/quality and token estimate. Repo, careful, new, writing, web, long context, visual, speed, ops/destructive/money, PR, verify, overnight, volume, local privacy and zero marginal cost are represented.
- Exhaustive account/model/effort ranking and factor contributions. Model/effort flags reach official CLIs. Account attribution, agy model pools, unknown quota, expiry, disabled CLIs, caps, paid opt-in and independent verifier constraints are explicit.
- Local is first-class for overnight/high-volume/zero-cost tasks. Privacy never falls back to cloud. Direct text endpoints are not represented as agentic repo editors. Fleet Ollama uses allowlisted SSH host names and loopback runtime, no new listener.
explain "<prompt>"and --json; pick/3 embeds full ranked/excluded explanation.- Native read-only Try a prompt, Routing map with 15 examples, Defaults and original account-health tab. Typed task/capability/quota/cost/rule pills, factors/runners-up, live policy/quota snapshot, manual Refresh and existing 300-second periodic refresh.
- DEC-RMR-35–38, README, INTERFACES, docs/ROUTING-V3.md, primary-source catalog snapshots and frozen 15-example golden fixture. Installed apps/config untouched.
Verification
| Check | Result | Artifact |
|---|---|---|
| Default swift test | 130 tests / 19 suites / 0 failures, rc 0 | tests-final.log |
| RMR_DISABLE_QUOTA_KIT=1 isolated scratch | 130 / 19 / 0, rc 0 | tests-noquota-final.log |
| Universal CLI + app build | rc 0, both x86_64 + arm64, Intel minos 26.7, plist valid | build-final.log |
| Release CLI acceptance | 9 cases, asserted expected return codes and constraints | release-acceptance.json |
| Full candidate factor arithmetic | Every release candidate score equals its factor sum | release-acceptance.json |
| Launch integration | Selected gpt-6-astra/high reaches plan and fake apply/audit | V3LaunchTests |
| Secret scan | 66 in-memory known values, zero findings; exact and pattern controls passed | secret-scan-precommit.json; postcommit follows |
| Whitespace | git diff --check rc 0 | branch diff |
| Native tab screenshots | All three captured and visually inspected, 1720×1480 | app-try.png, app-map.png, app-defaults.png |
| Artifact identity | SHA-256 for app, CLI, policy and screenshots | artifact-manifest.json |
Golden fixture: 2026-09-27T20:00:00Z; 33% fresh use for each account and zero observed CLI sessions. Fifteen prompts pin model, account, effort and full labels. Examples: careful → Codex gpt-6-astra/high; docs → Claude sonnet/medium; research → agy gemini-3.8-flash-high; quick → Codex gpt-6-luna/low; verify (builder codex) → Claude fable/high; bulk/quick local → Studio qwen3.5:4b-mlx.
Live release results differ legitimately: observed caps were agy 3/1, Claude 10/2, Codex 2/1 and Grok 1/1. Copilot auto won careful/verify among available agent CLIs; Hermes was a runner-up. Bulk and quick local chose Studio qwen3.5:4b-mlx; private summary chose Studio qwen3-coder-next:q4_K_M. Unknown builder and purchase returned 2/no route. Dry-run applied=false. Three paid candidates became eligible only with --allow-paid-api, in an explain call; no inference was run.
The GUI images are from the rebuilt worktree app, not mockups. Final observed process/window pairs: Try 34718/4588, Map 39374/4705, Defaults 49350/4748. Each owned process was terminated by exact PID/path. Native CUA returned a closed-pipe error; interactive clicking/pasting is NOT claimed. Read-only initial-tab settings plus LaunchServices supplied captures. A mislabeled intermediate Defaults capture was rejected during visual inspection and replaced with the correct tab.
Explicit limitations and unavailable discovery
- Numerical strength/speed/quality ratings are dated policy estimates, not measured benchmarks. Advertised context is not allocated runtime context. Peer context is inherited only when the full installed digest matches Studio, otherwise unknown.
- ToshLLM is installed on both Intel hosts with configured Qwen3.8 GGUFs. Port 8080 on rdmpw3265m and configured port 11435 on rdmpw3275m are not serving a verified model list. Earlier rdmpw3275m:8080 was OrbStack, not ToshLLM. The existing rdmodelrouter/TOSHLLM_API_KEY is present; values were pipe-only, never persisted. XEntropy binary is absent. No cloud URL/model/price was invented. Inventory is included and scored as unavailable; no ToshLLM inference adapter is claimed.
- AINetNode's source gateway port 11435 was unavailable on Studio/Intel probes. No models/services were started or downloaded to fabricate availability.
- Copilot exposes its native auto selector in help, not a verified concrete model list. Hermes model is its existing native config; runtime validates entitlement.
- Ollama accepts supplied text; the router does not enumerate millions of files or automatically provide repo contents. Batch extraction/orchestration is caller work.
- Fallback is ranked alternatives, not automatic replay of partially executed work.
- Topology audit command at the prescribed Documents path was absent (rc 127). Host identity and read-only SSH inventories were verified; no deployment or cross-account/session-store movement occurred. Preflight had historical HIGH mail; no outbound messages/acks/worker fan-out or force-claim was used.
Durability, undo and next action
Apple Notes PENDING: launchctl managername is Background. fleet-notes-publish says “In a headless run, write the file copy and report the Notes entry as PENDING”. Durable copies are under ~/dev/LLM/Claude/changelogs and ~/dev/LLM/Codex/changelogs. No Notes GUI workaround, installed replacement, signing, merge, push or paid call.
Resume: independently review the commit and this report, rerun the two test commands, and inspect the three screenshots. Lead owns integration/installed rollout. Undo source changes with a new revert commit. No user credential/model store was modified; earlier development bundles were archived by the build script.
Final commit receipt — 2026-09-27 17:19:55 EDT · rdmsm4x
Commit 6d5287ac8e47b3fed7aca20a4ff822ac1b1b8d1b on
rmr/routing-intel-20260927, Agent trailer present. 45
explicit files committed; worktree clean. No merge or push.
TASK-20260927-64 builder work resolved; lead review/integration remains
separate. All owned development app processes closed. Postcommit secret
scan: 66 in-memory known values, 177 files, 326 reachable blobs and
current transcript; zero findings, both positive controls passed. Notes
helper returned rc 3 (Background rather than Aqua); both changelog
copies retained.